Journal article
Multi-granularity 3D Visual Language Tracking from monocular videos: A benchmark and model
Hongkai Wei, Rong Wang, Keyu Guo, Yongle Huang, Shijie Sun, Xiangyu Song, Mingtao Feng, Naveed Akhtar
Pattern Recognition | Elsevier BV | Published : 2027
Abstract
Object tracking, a fundamental task in computer vision, has recently evolved beyond purely visual cues to incorporate natural language as high-level semantic guidance. Monocular 3D Visual Language Tracking (Mono3DVLT) extends this paradigm into the 3D domain, seeking to localize and track a referred 3D object in monocular video under natural language guidance. Despite its potential, existing approaches remain constrained by single-granularity language descriptions and weak temporal reasoning. Motivated by the diversity of natural instructions in real-world traffic scenes, we introduce a comprehensive framework for evaluating whether monocular 3D trackers perform under controlled variations i..
View full abstractGrants
Awarded by National Natural Science Foundation of China
Awarded by National Key Research and Development Program of China